# CEO review of the PRD Oct 1, 2026 · reviewed with Claude Code (`/plan-ceo-review`, condensed) The PRD points the right way and is honest about its own weak spots. Its problem is where the effort goes: the demo spends most on the part of the product that is becoming free, and plans to cut first the one piece that makes a statistician believe the numbers. The owner accepted all six proposals on 1 Oct 2026, and they are applied to `PRD.md`, `demo-architecture.md`, `CLAUDE.md` and `BACKLOG.md`. ## How this review ran The interactive question flow was skipped at the owner's request. Two process choices were made without asking, and both are easy to reverse: - **No `/office-hours` pass.** The PRD and `prd-review.md` already hold the problem statement, premises and six reader critiques. - **Mode: hold scope.** No feature is cut or added here. Reallocations are proposed for the owner to decide (P-1 to P-6 below). Market context comes from PRD Appendix D, researched the same day. No new web search was run. ## Verdict | Question | Answer | | --- | --- | | Is the problem real? | Yes. Designs built on Western assumptions, and lessons lost at study close, are both documented in Appendix A. | | Does the demo test the risky premise? | Partly. It tests whether a study *can be ingested*. It does not yet test whether the profile's numbers *can be believed*, which is the question every buyer in §6 is really asking. | | Is the scope buildable? | Only with zero slack, as the plan itself says. Real data is not confirmed until week 2. | | What would most improve it? | Move effort from document extraction to calibration and replay, and give the "global defaults" baseline a real source. | ## Findings ### 1. Effort goes to the commodity part The document pipeline, retrospective and feedback digest take 14 of 48 engineer-weeks. Appendix D names protocol extraction and drafting as turning into a free feature of model vendors (Claude for Life Sciences, Faro, Veeva). The parts nobody else has are the behaviour profile and a simulator that knows how sites behave. Replay is the only demo step that answers the biostatistician's question ("no unexplained result under questioning"), and it is listed first to drop. ### 2. The headline comparison has no baseline source Act 4, "Simulate twice", compares the local profile against global defaults. Every `global_default` in `schema/behaviour-parameter.example.json` has `"source": "placeholder"`. If the baseline is unsourced, the difference the audience sees is unsourced too. Appendix A already holds published values: dropout 19.1%, screen failure 36.3%, site activation 31.4 weeks. ### 3. The synthetic package can be the trust argument, not just a fixture The generator knows the true parameters. Regenerate the study 200 times from different seeds, estimate each, and count how often the stated 90% interval holds the truth. That is a calibration result the team can show before any real data arrives, and it is cheap once the generator and estimator exist. It also protects the demo if the week-2 real-data decision is no-go. ### 4. Intervals from one study must say what they are A prediction interval for a *new* study needs the spread between studies, which one study cannot measure. The plan says this. The screens do not show it yet: the evidence table on the Report screen has no interval column, and "Ethics review rounds 2 · Assumed" is a bare number. Under hard rule 4 that is a defect. ### 5. Two servers are not needed until real data exists On synthetic data, the hospital zone and platform zone can run as two processes on one machine, with the boundary enforced by tests: no platform import from the hospital zone, separate databases, and one export file as the only crossing. The hospital server request (backlog 0.8) still goes in week 1, because approval is slow. ### 6. The demo ending is still open `BACKLOG.md` records that ending on the design report versus a dossier draft is unanswered. The design report is the first thing a buyer pays for (§8), and the dossier is out of scope (F6). The report is the right ending. ## Proposed changes (accepted 1 Oct 2026) | ID | Proposal | Effort | Status | | --- | --- | --- | --- | | P-1 | Make replay (F3.8) protected, not first to drop. Drop scanned-document text recognition from the demo instead: accept born-digital Word and PDF, and label scans "not read in the demo". | M | Accepted; applied | | P-2 | Add a sourced global-defaults table. Every default carries its citation, and any default without one is labelled `assumed` and flagged on screen. | S | Accepted; applied | | P-3 | Add a calibration check on the synthetic package: coverage of stated intervals across repeated seeds, shown as a "Can these intervals be trusted?" panel in act 5. | S | Accepted; applied | | P-4 | Run both zones on one machine as two processes until real data arrives. Keep the boundary in tests. | S | Accepted; applied | | P-5 | End the demo on the signed design report, not a dossier draft. | S | Accepted; applied | | P-6 | Add an interval column to the evidence table on the Report screen, and show the between-study spread source on every prediction interval. | S | Accepted; applied | ## Error and rescue map for the golden path | Step | What can fail | Class | What the user sees | Test | | --- | --- | --- | --- | --- | | Intake | A file of unknown type | `UnknownFileType` | Listed as "Unrecognised" on the checklist; the path continues | file typing test | | De-identification | A direct-identifier column or a phone number in free text survives | `DeidentificationError` | Intake stops; the log names the column | de-identification test | | Consent | Dataset has no consent scope | `ConsentScopeMissing` | Estimator refuses to run | consent test | | Export | A cell below the minimum size, a patient-level field, or no approver | `ExportRejected` | Export blocked with the reason; nothing crosses | export gate tests | | Profile store | `learned` with under 30 patients and under 3 sites | `LabelRuleViolation` | Save refused | label rule test | | Clinician review | A decision without a named reviewer | `ReviewError` | Decision refused | review test | | Simulation | Run without seed, profile version or code version | `ReproducibilityError` | Run refused | reproducibility test | | Replay | A record dated after the cut reaches the fit | `LeakageError` | Replay stops | leakage test | | Report | Export without a named signer, or a number without a label | `SignOffRequired`, `UnlabelledNumber` | Export blocked | report tests | | Gateway | Input tagged patient-level | `PatientLevelBlocked` | Call refused and logged | gateway test | No catch-all handlers. Every failure stops its step, names the reason and leaves an audit entry. ## Failure modes worth naming - **The profile looks precise and is wrong.** One study gives narrow within-study intervals. Rescue: show the elicited between-study spread on every prediction interval, and the calibration panel (P-3). - **The baseline is made up.** Rescue: P-2. - **Real data never arrives.** Rescue: the synthetic path is complete and honest about being synthetic (hard rule 8), and the PRD's re-framing rule in `prd-review.md` applies. - **The demo is "80% done everywhere".** Rescue: the golden path runs end to end first, then each step deepens. ## Not in scope As already listed in `CLAUDE.md`: F7, F6, F3.1, F4.2, F8.1 beyond a schema check, F1.7, F9.9. Nothing was added. ## What is being built now The thin golden path on the synthetic package, as `CLAUDE.md` asks. This changes no PRD scope. The plan is in `docs/build-plan.md`. ## GSTACK REVIEW REPORT | Review | Runs | Status | Findings | | --- | --- | --- | --- | | CEO review (`/plan-ceo-review`) | 1 | DONE_WITH_CONCERNS | 6 findings, 6 proposals pending owner decision | - Mode: HOLD SCOPE, chosen without asking because the owner declined the question flow. - Outside voice: not run. - VERDICT: sound direction; effort should move from extraction to calibration and replay. - Follow-up: the thin build confirmed finding 3. The calibration check found three interval-method problems and the replay test found a leakage bug in its first hour of use. **UNRESOLVED DECISIONS:** - Open decisions carried from `CLAUDE.md`: which study, hospital hosting, local language model, front-end framework, lead indication